The 1997 Abbot System for the Transcription of Broadcast News
نویسندگان
چکیده
This paper describes the development of a connectionist-hidden Markov model (HMM) system for the 1997 DARPA Hub-4E CSR evaluations. We describe both system development and the enhancements designed to improve performance on broadcast news data. Both multilayer perceptron (MLP) and recurrent neural network acoustic models have been investigated. We assess the effect of using gender-dependent acoustic models, and the impact on performance of varying both the number of parameters and the amount of training data used for acoustic modelling. The use of contextdependent phone models is described, and the effect of the number of context classes is investigated. We also describe a method for incorporating syllable boundary information during search. Results are reported on the 1997 DARPA Hub-4E development test set. We then describe the CU-CON evaluation system and report results on the 1997 Hub-4E test set.
منابع مشابه
Transcription of broadcast television and radio news: the 1996 ABBOT system
This paper describes the development of the cu-con system which participated in the 1996 ARPA Hub 4 Evaluations. The system is based on Abbot, a hybrid connectionist-HMM large vocabulary continuous speech recognition system developed at the Cambridge University Engineering Department [4]. The Hub 4 Evaluation task involves the transcription of broadcast television and radio news programmes. Thi...
متن کاملTranscription of Broadcast Television and Radio News : The
This paper describes the development of the cu-con system which participated in the 1996 ARPA Hub 4 Evaluations. The system is based on Abbot, a hybrid connec-tionist-HMM large vocabulary continuous speech recognition system developed at the Cambridge University Engineering Department 4]. The Hub 4 Evaluation task involves the transcription of broadcast television and radio news programmes. Thi...
متن کاملTranscribing broadcast news with the 1997 Abbot System
Recent DARPA CSR evaluations have focused on the transcription of broadcast news from both television and radio programmes [17]. This is a challenging task because the data includes a variety of speaking styles and channel conditions. This paper describes the development of a connectionist-hidden Markov model (HMM) system, and the enhancements designed to improve performance on broadcast news d...
متن کاملThe THISL Spoken Document Retrieval System
THISL is an ESPRIT Long Term Research Project focused the development and construction of a system to items from an archive of television and radio news broadcasts. In this paper we outline our spoken document retrieval system based on the ABBOT speech recognizer and a text retrieval system based on Okapi term-weighting . The system has been evaluated as part of the TREC-6 and TREC-7 spoken doc...
متن کاملTranscription of broadcast news-system robustness issues and adaptation techniques
This paper describes some of the main problems and issues speci c to the transcription of broadcast news and describes some of the methods for solving them that have been incorporated into the IBM Large Vocabulary Continuous Speech Recognition System
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1998